Papers with sentence representation models
UNSEE: Unsupervised Non-contrastive Sentence Embeddings (2024.eacl-long)
Copied to clipboard
| Challenge: | Unsupervised Non-Contrastive Sentence Embeddings demonstrates better performance compared to SimCSE in the Massive Text Embing (MTEB) benchmark. |
| Approach: | They introduce UNSEE, which stands for Unsupervised Non-Contrastive Sentence Embeddings, which demonstrates better performance compared to SimCSE in the Massive Text Embing benchmark. |
| Outcome: | The proposed solution achieves better performance than contrastive objectives on the Massive Text Embedding (MTEB) benchmark. |
Adaptive Reinforcement Tuning Language Models as Hard Data Generators for Sentence Representation (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods use contrastive learning (CL) to learn effective sentence representations, but require extensive human annotation. |
| Approach: | They propose a reinforcement learning approach for fine-tuning small-parameter LLMs to generate high-quality hard contrastive data without human feedback. |
| Outcome: | The proposed method achieves state-of-the-art on seven semantic text similarity tasks. |
Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence Reasoning (2023.eacl-main)
Copied to clipboard
| Challenge: | Recent studies show that textual entailment learning reduces social biases in pretrained sentence encoders. |
| Approach: | They compare pretrained sentence encoders with textual entailment models that learn language logic for downstream language understanding tasks. |
| Outcome: | The proposed models outperform models with lower bias without debiasing processes on stereotype, profession, and emotion bias tests. |
Investigating BERT’s Knowledge of Language: Five Analysis Methods with NPIs (D19-1)
Copied to clipboard
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman
| Challenge: | Recent work evaluating sentence representation models' knowledge of grammar has been slower to emerge. |
| Approach: | They propose five experimental methods inspired by prior work evaluating pretrained sentence representation models to examine their grammatical knowledge. |
| Outcome: | The proposed methods show that the model has significant knowledge of the licensing environment but its success varies widely across different methods. |
Evaluating Multilingual Sentence Representation Models in a Real Case Scenario (2022.lrec-1)
Copied to clipboard
| Challenge: | a recent study has shown that the infamous Protocols are actually plagiarized . a convoluted task with no standard benchmarks for paraphrase detection and sentence similarity is a problem . |
| Approach: | They evaluate sentence representation models on the paraphrase detection task . they use a forged text from the so-called "Protocols of the Elders of Zion" scholars have demonstrated that the first text plagiarizes from the second . |
| Outcome: | The proposed model is based on the forged “Protocols of the Elders of Zion” . the model is similar to the standard model but has some problems . |